Papers with vision-aided unsupervised constituency parsing

1 papers
Vision-aided Unsupervised Constituency Parsing with Multi-MLLM Debating (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches require explicit cross-modal alignment, but new approaches address these challenges.
Approach: They propose a framework for vision-aided unsupervised constituency parsing . they leverage multimodal large language models pre-trained on diverse image-text or video-text data .
Outcome: The proposed framework achieves state-of-the-art performance on image-text and video-text datasets, improving robustness and accuracy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations